Papers with machine learning technologies
Self-supervised Representation Learning for Speech Processing (2022.naacl-tutorials)
Copied to clipboard
Hung-yi Lee, Abdelrahman Mohamed, Shinji Watanabe, Tara Sainath, Karen Livescu, Shang-Wen Li, Shu-wen Yang, Katrin Kirchhoff
| Challenge: | Self-supervised representation learning (SSL) uses proxy supervised learning tasks to obtain training data from unlabeled corpora. |
| Approach: | They propose to survey the latest SSL techniques, tools, datasets, and performance achievement in speech processing to scale up current machine learning technologies. |
| Outcome: | The proposed tutorial is highly relevant to the special theme of ACL about language diversity. |
The D-WISE Tool Suite: Multi-Modal Machine-Learning-Powered Tools Supporting and Enhancing Digital Discourse Analysis (2023.acl-demo)
Copied to clipboard
| Challenge: | The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded. |
| Approach: | They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH) |
| Outcome: | The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research. |
Designing Multilingual Interactive Agents using Small Dialogue Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | a new study aims to develop a design framework for multilingual interactive agents . large amounts of data and language resources are needed to develop most key components . |
| Approach: | They propose a general design framework for multilingual interactive agents in specialized domains with small or non-existent dialogue corpora. |
| Outcome: | The proposed framework integrates external language services for supporting multilingual functions and realizes context-aware dialogue generation under the situation of small corpora. |
MDACE: MIMIC Documents Annotated with Code Evidence (2023.acl-long)
Copied to clipboard
Hua Cheng, Rana Jafari, April Russell, Russell Klopfer, Edmond Lu, Benjamin Striner, Matthew Gormley
| Challenge: | Computer-Assisted Coding (CAC) systems are required to provide supporting textual evidence to justify billing codes. |
| Approach: | They propose a dataset for evidence/rationale extraction on an extreme multi-label classification task over long medical documents. |
| Outcome: | The proposed dataset can be used to evaluate evidence extraction methods for CAC systems, as well as the accuracy and interpretability of deep learning models for multi-label classification. |